Papers by Marta Gonzalez Mallo
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering (2025.naacl-short)
Copied to clipboard
Anna Arias-Duart, Pablo Agustin Martin-Torres, Daniel Hinjos, Pablo Bernabeu-Perez, Lucia Urcelay Ganzabal, Marta Gonzalez Mallo, Ashwin Kumar Gururajan, Enrique Lopez-Cuena, Sergio Alvarez-Napagao, Dario Garcia-Gasulla
| Challenge: | Current Large Language Models (LLMs) benchmarks are often based on open-ended or close-ended QA evaluations, avoiding the requirement of human labor. |
| Approach: | They propose a multi-axis suite for healthcare LLM evaluation, exploring correlations between open and close benchmarks and metrics. |
| Outcome: | The proposed framework explores correlations between open and close benchmarks and metrics in the healthcare domain, with blind spots and overlaps in existing methodologies. |